Skip to main content

5-Gradients

5- Understanding Gradients​

What is a Gradient?​

Gradient is simply relationship of weight and loss.

if Increasing w tends to increase loss -> negative gradient. -> we need to decrease weight

if Increasing w tends to decrease loss -> positive gradient -> we need to increase weight

Calculating Gradients Manually​

Let's calculate the exact gradient of our loss function and use it to update the weight and bias. parameternew​=parameter−learning rate×gradient

Gradient Descent (Weight and Bias Updation)

Why does this work?​

By calculating the derivatives (`dw` and `db`), we figure out exactly how a tiny change in `w` or `b` affects the overall loss. Multiplying these gradients by a small learning rate ensures we take careful, controlled steps towards the minimum possible error!

Optimization​

Lets see how convergence based optimization works, basically the farther you are from answer, the bigger steps you take to come near answer, while as you get close you take smaller steps why? because the step amount we decide based on steepness of the gradient. steeper it is, means our answer is very wrong. so we simply take a bigger step. lets see how its done -

dw = 2 * (prediction - target) * x

this is dw/dl the small change there should be in w based on small change in loss. if loss is big, this (prediction - target) value will be also big.

and then we simply do the actual weight change using learning rate, if loss big-> dw big-> w update big. if loss small -> dw small -> w update small.

w = w - (learning_rate * dw)